Regression versus classification
The first question to ask when you start a machine learning project is: what kind of thing am I trying to predict? If the answer is a label or category, you reach for a classifier. If the answer is a number somewhere on a continuous scale, you reach for a regressor.
- House price ($312,000)
- Temperature tomorrow (23.4 °C)
- Patient's blood pressure (128 mmHg)
- Time a delivery will take (47 minutes)
- Number of sales next quarter (8,420 units)
- Will this loan default? (Yes / No)
- What species is this? (Cat / Dog / Bird)
- Is this email spam? (Spam / Not spam)
- What digit is this? (0 through 9)
- Is this tumour benign? (Benign / Malignant)
The boundary can get fuzzy. You can turn a regression problem into a classification one by bucketing the output: instead of predicting the exact house price, predict whether it is "cheap," "mid-range," or "expensive." But for anything where the exact number matters, regression is the right tool.
Linear regression: fit a line, predict a number
Linear regression is the simplest regression algorithm, and it is genuinely useful. The core idea is that you draw a straight line through your data in a way that minimises the total distance between the line and all your data points. Future predictions are just reading off the line at the relevant x-value.
You are studying for an exam and you have records from previous students: how many hours each person studied and what score they got. You plot this as a scatter of dots. Linear regression draws the single best straight line through that scatter. Once the line is drawn, you can read off "if I study 7 hours, I should expect a score of about 74."
The model's job during training is to find the values of m and b that make the line fit the training data as well as possible. These are the parameters of the model, the numbers the model learns. Once training is done, m and b are fixed. Predicting is then just plugging in any x and computing the result.
When you have multiple input features, the line becomes a hyperplane, but the principle is identical. Instead of one slope m, you have one weight (slope) for each feature. Still the same core idea: find the weights that minimise prediction error.
The green line is the best fit through the data. The orange dashed lines are residuals: the vertical distance between each point and the line. Linear regression finds the line that minimises the total size of these residuals.
Loss functions: measuring how wrong you are
To fit the line, the model needs a way to measure how wrong its current parameters are. That measure is the loss function. In regression, the most common approach is to look at the residuals: the gap between each prediction and the actual value. The loss function summarises all those gaps into a single number.
Three metrics dominate regression evaluation. Understanding what each one rewards and what each one punishes will save you from misinterpreting your model's performance.
Use MAE when you want an error in the same units as the target (dollars, degrees, minutes) and need to explain the result to a non-technical audience. Use MSE or RMSE (root of MSE) during model training and comparison. Report R-squared alongside MAE when presenting to stakeholders: "The model explains 85% of price variation and is off by an average of $12,000."
When a straight line is not enough
Linear regression assumes the relationship between input and output is a straight line. Real-world relationships often curve. A student's scores might improve steadily with study hours up to a point, then plateau due to fatigue. A car's fuel efficiency might drop slowly at moderate speeds then plummet at highway speeds.
Polynomial regression handles this by adding powers of the feature as extra columns. Instead of just using x, you also include x squared, x cubed, and so on. The model is still linear in its parameters (it learns weights for each term) but the resulting curve can bend and flex to fit non-linear patterns.
More complex relationships call for more powerful algorithms. Decision trees can do regression (predict numbers at each leaf instead of class labels). Random Forests average the predictions of many trees. Gradient Boosting builds trees sequentially, each one correcting the errors of the last. These ensemble methods consistently rank among the best performers on tabular data with complex, non-linear relationships.
Predicting house prices in Scikit-learn
The classic regression toy problem is predicting house prices from features like size, location, number of rooms, and age. Here is a full pipeline using the California housing dataset that ships with Scikit-learn.
from sklearn.datasets import fetch_california_housing from sklearn.linear_model import LinearRegression from sklearn.model_selection import train_test_split from sklearn.metrics import mean_absolute_error, r2_score import numpy as np # Load the built-in California housing dataset data = fetch_california_housing() X, y = data.data, data.target # Split into training and test sets X_train, X_test, y_train, y_test = train_test_split( X, y, test_size=0.2, random_state=42 ) # Train a linear regression model model = LinearRegression() model.fit(X_train, y_train) # Predict on the test set predictions = model.predict(X_test) # Evaluate: MAE is in units of $100,000 (dataset scale) mae = mean_absolute_error(y_test, predictions) r2 = r2_score(y_test, predictions) print(f"MAE: ${mae * 100_000:,.0f}") print(f"R-squared: {r2:.3f}") # Check the learned weights for each feature for name, coef in zip(data.feature_names, model.coef_): print(f" {name}: {coef:.4f}")
R-squared: 0.576
MedInc: 0.4487
HouseAge: 0.0099
AveRooms: -0.1073
AveBedrms: 0.6451
Population: -0.0000
AveOccup: -0.0038
Latitude: -0.4214
Longitude: -0.4333
The model is off by about $52,000 on average and explains 57.6% of the variation in house prices. For a linear model with no feature engineering, that is a reasonable starting point. The coefficients reveal something important: median income (MedInc) has the largest positive weight, meaning higher income areas are the strongest predictor of higher house prices. Latitude and longitude have large negative weights, encoding the fact that houses in coastal California tend to cost more.
You could improve this significantly. Polynomial features, better feature engineering, or switching to a gradient-boosted tree would likely get R-squared above 0.85. But for understanding the core idea, a simple linear model is always the right first step. Understand what a simple model says before reaching for something complex.
"All models are wrong, but some are useful."
George Box, statistician